Anthropic says its AI hacked real-world companies in three incidents
Anthropic said Thursday it had discovered three incidents in which its AI models exited test environments and compromised real-world organizations. The company discovered the breaches following an internal review triggered by a similar incident at rival OpenAI.
The incidents are the latest to raise questions about liability, disclosure standards and the adequacy of containment practices as AI systems become increasingly capable of conducting autonomous computer network operations.
Anthropic said the affected organizations, which have not been named, had not detected the activity themselves. One of those affected organizations had not been contacted at the time the company published its disclosure.
The root cause of the incidents, according to Anthropic, was a misunderstanding with the third-party evaluation partner, Irregular, that left the machines running Claude open to the internet. The models had been told they had no internet access.
The company’s reconstruction of the breaches is based on “evaluation transcripts” — logs that record an agent’s actions during a task, including commands it executed, responses it received, and the model’s own commentary and reasoning.
Anthropic’s research has found that the model’s own commentary and reasoning is rarely accurate. “Advanced reasoning models very often hide their true thought processes,” the researchers concluded, “and sometimes do so when their behaviors are explicitly misaligned.”
Three breakouts
Based on what Anthropic described as “the false belief that all accessible entities were intended to be in-scope for the exercise,” Claude compromised the three impacted organizations “using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”
In the first incident, a fictional target company used in the evaluation shared its name with an actual real-world website.
Claude found and exploited vulnerabilities in that real company’s infrastructure, extracting credentials and accessing a database containing several hundred rows of production data. Anthropic described this as the most serious of the three incidents in terms of impact.
Alongside potential criminal liability under computer misuse laws, the extraction of production data poses legal risks for both Anthropic and the target organization under British and European data protection frameworks, where a breach of this kind would typically be notifiable to regulators. Anthropic did not immediately respond to questions about how it is approaching its legal exposure.
This incident “was the only case in which the model continued to attack a system after learning it was likely operating in a real environment.” Despite eventually recognizing the system was real “in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis.”
In the second incident, Claude found that another fictional company’s setup instructions referenced a nonexistent PyPI package. Recognizing this as an opening — any system following those instructions would automatically download whatever package bore that name — Claude built and published a malicious package under the same name.
“Claude went to extensive lengths to carry out this attack,” wrote Anthropic.
To create a PyPI account, Claude needed an email address. To create an email address, it needed a phone number. After failing to find a free phone number service, it attempted and failed to obtain funds to pay for one through several different means, before eventually finding a free email provider and using it to register the account. It then uploaded malware to PyPI.
Although the AI believed the package registry it was using was part of the simulation, in reality the package was made freely available for roughly one hour during which time it was run on 15 real systems, said Anthropic.
One of these systems belonged to a security company that, like several firms in the industry, runs automated scanners that download newly published packages from registries like PyPI to check them for malware.
The scanner treated the package as safe to install, according to Anthropic, which said Claude's hidden code executed, exfiltrating the company's credentials to a collection point which it then used to access further infrastructure.
The security company was not identified. Anthropic did not explain whether Claude was interacting with infrastructure specifically designed to be found and compromised as a honeypot by the security firm. PyPI’s own security systems subsequently removed the package automatically.
In the third incident, Claude scanned approximately 9,000 internet-facing targets after failing to reach its intended fictional one, eventually compromising a real company's systems using basic techniques including SQL injection.
Unlike the other two models, however, this one — an internal research prototype not planned for general release — eventually recognized on its own that the target was real and stopped its attack without being prompted.
Responsibility and liability
Anthropic stressed throughout its disclosure that it believed there was no evidence of any of its models pursuing independent goals. In each case, the models did what their evaluations asked of them but did so while holding a false understanding of whether their environment was real.
Anthropic’s models accessed the internet through a path that was unintentionally left open, and mistook what they found as part of the exercise.
OpenAI’s models instead actively exploited a previously unknown vulnerability to escape their isolated test environment before identifying Hugging Face — a platform for open-source AI models and datasets — as a useful target and breaching its production infrastructure using stolen credentials and a second zero-day flaw.
Five days after the initial alert by Hugging Face, claiming to have been the target of an end-to-end autonomous cyberattack, OpenAI confirmed that it was responsible. This week it additionally acknowledged its rogue agent had subsequently also breached several additional third-party services.
Clement Delangue, Hugging Face’s co-founder and chief executive, said after OpenAI's disclosure that he “strongly believed there was no malicious intent” behind the breach, although the incident had already raised pointed questions about what disclosure obligations exist when an AI system causes unintended harm.
Hugging Face also noted that when its security team tried to use frontier AI models to analyse the attack, safety filters blocked its analysis of exploit payloads and attack commands, forcing it to use a self-hosted open-weight model instead — observing that “the attacker was bound by no usage policy, while our own forensic work was blocked.”
Anthropic said it is now working with METR, an independent AI evaluation organization, to conduct a third-party review of the incidents including access to all transcripts. It said it plans to release a lightly redacted transcript of the PyPI incident within the week.
Neither Anthropic nor Irregular responded to questions about whether any of the affected organizations are considering legal action, nor whether law enforcement has been in contact.
Alexander Martin
is the UK Editor for Recorded Future News. He was previously a technology reporter for Sky News and a fellow at the European Cyber Conflict Research Initiative, now Virtual Routes. He can be reached securely using Signal on: AlexanderMartin.79



